Papers with linguistic theory

19 papers
What Can String Probability Tell Us About Grammaticality? (2026.tacl-1)

Copied to clipboard

Challenge: linguistic theories have argued that language models have largely achieved grammatical competence, but they will assign non-zero probability to all strings.
Approach: They propose a theoretical framework for analyzing string probabilities in linguistics based on simple assumptions about the generative process of corpus data.
Outcome: The proposed framework makes three predictions using 280K sentence pairs in English and Chinese.
Joint Universal Syntactic and Semantic Parsing (2021.tacl-1)

Copied to clipboard

Challenge: Several attempts have been made to jointly parse syntax and semantics, but this trade-off is not well understood.
Approach: They propose multiple model architectures that exploit the rich syntactic and semantic annotations contained in the Universal Decompositional Semantics dataset to obtain state-of-the-art results.
Outcome: The proposed model outperforms existing models in 8 languages and their results are consistent across languages.
The Importance of Modeling Social Factors of Language: Theory and Practice (2021.naacl-main)

Copied to clipboard

Challenge: Current NLP models focus on information content while ignoring language’s social factors.
Approach: They propose that NLP systems focus on information content while ignoring language’s social factors to improve performance.
Outcome: The proposed approach improves the performance of existing systems, open up new applications, and increase fairness and usability for all users.
Linguistically Grounded Analysis of Language Models using Shapley Head Values (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods for probing language models for morphosyntactic constructions are not well understood . language models gain knowledge of grammatical phenomena during pretraining, but exactly how this knowledge is encoded is not well established.
Approach: They propose a method for probing language models via Shapley Head Values . they use a BLiMP dataset to test their method on linguistic constructions based on a Shaply Head Value method .
Outcome: The proposed method can be used to investigate linguistic knowledge in language models . it shows that attention heads responsible for processing related linguistic phenomena cluster together .
Dead or Murdered? Predicting Responsibility Perception in Femicide News Reports (2022.aacl-main)

Copied to clipboard

Challenge: linguistic expressions of gender-based violence can conceptualize the same event from different perspectives by emphasizing certain participants over others.
Approach: They conduct a large-scale perception survey of GBV descriptions from italian newspapers and train regression models that predict the salience of GV participants with respect to different dimensions of perceived responsibility.
Outcome: The proposed model shows that salient focus is more predictable than salient blame, and perpetrators’ salience is more predictable than victims’ salient.
Abstract Meaning Representation for Multi-Document Summarization (C18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is a semantic representation of natural language based on linguistic theory .
Approach: They propose to use Abstract Meaning Representation (AMR) as a content representation.
Outcome: The proposed framework is fully data-driven and flexible.
Debiasing Word Embeddings with Nonlinear Geometry (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for debiasing word embeddings are limited to individual social categories . however, real-world corpora typically present multiple social categories that may correlate or intersect with each other.
Approach: They propose a method to debias word embeddings using nonlinear geometry of individual biases.
Outcome: Empirical results show that the proposed method mitigates biases associated with individual social categories and treats each category in isolation.
No Questions are Stupid, but some are Poorly Posed: Understanding Poorly-Posed Information-Seeking Questions (2025.acl-long)

Copied to clipboard

Challenge: When a question is poorly posed, answerers struggle to converge on dominant interpretations, while models attempt comprehensive coverage by addressing many interpretations simultaneously.
Approach: They propose a computational framework to study poorly-posedness of questions by generating spaces of potential interpretations and computing distributions based on interpretations chosen by answerers in the Reddit question thread.
Outcome: The proposed framework analyzes poorly-posed questions using a set of interpretations chosen by human answerers and large language models.
Language Modelling as a Multi-Task Problem (2021.eacl-main)

Copied to clipboard

Challenge: Using multitask learning, humans are optimising their behaviour towards a multitude of objectives to reach their goals in dayto-day life.
Approach: They propose to study language modelling as a multi-task problem by examining the generalisation behaviour of language models as they learn the linguistic concept of Negative Polarity Items.
Outcome: The proposed model is able to learn the linguistic concept of Negative Polarity Items (NPIs) and is a multi-task learning model.
Penguins Don’t Fly: Reasoning about Generics through Instantiations and Exceptions (2023.eacl-main)

Copied to clipboard

Challenge: Generics express generalizations about the world that are not universally true . commonsense knowledge bases encode some generic knowledge but rarely enumerate exceptions .
Approach: They propose a framework informed by linguistic theory to generate exemplars for generics . they generate 19k exemplar cases for 650 generics and show they outperform a strong baseline .
Outcome: The proposed framework outperforms a baseline framework by 12.8 precision points.
Variance of Average Surprisal: A Better Predictor for Quality of Grammar from Unsupervised PCFG Induction (P19-1)

Copied to clipboard

Challenge: In unsupervised grammar induction, data likelihood is only weakly correlated with parsing accuracy, especially at convergence after multiple runs.
Approach: They propose to use VAS instead of data likelihood to find better grammars by examining linguistically-motivated constraints related to syntax.
Outcome: The proposed model is better suited for word order typology classification than data likelihood.
What company do words keep? Revisiting the distributional semantics of J.R. Firth & Zellig Harris (2022.naacl-main)

Copied to clipboard

Challenge: linguists J.R. Firth and Zellig Harris are often credited with the invention of "distributional semantics" a close reading of their work uncovers two distinct and in many ways divergent theories of meaning .
Approach: They propose to compare two different theories of meaning that focus on internal workings of linguistic forms with a broader cultural and situational context.
Outcome: The authors examine the differences between their theories of meaning and the internal workings of linguistic forms . they find that Firth could guide the field towards a more culturally grounded notion of semantics .
CoPrUS: Consistency Preserving Utterance Synthesis towards more realistic benchmark dialogues (2025.coling-main)

Copied to clipboard

Challenge: Large-scale Wizard-Of-Oz dialogue datasets lack certain types of utterances, which would make them more realistic.
Approach: They propose to use a large language model to create and repair communication errors in an automatic pipeline.
Outcome: The proposed method is based on linguistic theory and uses a state-of-the-art Large Language Model (LLM) to create the error and repair it.
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: a new analysis leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs).
Approach: They propose to use linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs).
Outcome: The proposed analysis reveals that linguistic similarity is significantly influenced by training data exposure, leading to higher cross-LLM agreement in higher-resource languages.
Language-specific Effects on Automatic Speech Recognition Errors for World Englishes (2022.coling-1)

Copied to clipboard

Challenge: Existing systems are not able to meet the needs of speakers of different demographic groups.
Approach: They propose to analyze the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors.
Outcome: The proposed system predicts certain errors from the phonological structure of a speaker’s native language.
Revisiting Supertagging for faster HPSG parsing (2024.emnlp-main)

Copied to clipboard

Challenge: a new supertagger for HPSG-based treebanks is used to improve parsing speed and accuracy.
Approach: They propose to integrate the best supertagger into an HPSG-based parser and compare it to an existing system.
Outcome: The proposed system achieves 97.26% accuracy on 950 sentences from WSJ23 and 93.88% on the out-of-domain technical essay The Cathedral and the Bazaar.
Causal Interventions Reveal Shared Structure Across English Filler–Gap Constructions (2025.emnlp-main)

Copied to clipboard

Challenge: Language Models (LMs) have emerged as powerful sources of evidence for linguists seeking to develop theories of syntax.
Approach: They propose to use causal interpretability methods to characterize abstract mechanisms that LMs learn to use by transferring a wh-filler-gap structure into a gap-less c++ class.
Outcome: The proposed methods can characterize the abstract mechanisms that LMs learn to use, and challenge claims that they can be learned only with strong innate priors.
Specifying Genericity through Inclusiveness and Abstractness Continuous Scales (2024.lrec-main)

Copied to clipboard

Challenge: Using a pilot study, we created a small but crucial annotated dataset of 324 sentences, demonstrating the framework’s effectiveness in capturing nuanced aspects of genericity.
Approach: They propose a framework for fine-grained modeling of noun phrases' genericity in natural language using a small but crucial annotated dataset of 324 sentences.
Outcome: The proposed framework can be used to model genericity of noun phrases in natural language and can be easily compared with existing binary annotations.
Do Language Models Use Logophoric Cues? Evidence from Mandarin Chinese Long-Distance Reflexive (2026.findings-acl)

Copied to clipboard

Challenge: Using minimal pairs and surprisal-based measures, we assess whether large language models exhibit systematic biases toward non-local antecedents in logophoric contexts.
Approach: They examine large language models’ sensitivity to four logophoric cues known to license long-distance binding of the reflexive ziji .
Outcome: The proposed model families show that they exhibit above-chance sensitivity to all four cues, while lexically anchored cue are more robustly captured than discourse-level cue.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations